In this lesson
The Future of AI: What Is Coming Next
We are at an inflection point. The systems that exist today are already remarkable. The systems being built right now will be more capable still. Understanding where the field is heading equips you to make better decisions, build more responsibly, and position yourself well for what comes next.
Looking Ahead
Predicting the future of a fast-moving technology field is difficult, and anyone who claims certainty about where AI will be in five or ten years should be viewed sceptically. What we can do is identify the trajectories that are already under way, understand the forces driving them, and think carefully about their implications.
Several themes are shaping the near-term future of AI. Models are becoming multimodal, able to process and generate text, images, audio, and video in an integrated way. Systems are gaining the ability to take actions in the world through tools and external APIs, moving from passive responders to active agents. Reasoning capabilities are improving dramatically, with dedicated architectures trained specifically to think through problems step by step. The ecosystem is splitting along an open versus closed axis. And the field is grappling seriously with safety and alignment as capabilities advance.
This final lesson surveys each of these areas, grounding the discussion in research and systems that already exist, while being honest about what remains uncertain.
Earlier lessons built conceptual foundations with well-established techniques. This lesson deals with an area of active development where understanding is still forming. Some of what is described here as "near-term" may already have changed by the time you read it. Treat this as a structured orientation to a living conversation rather than a static reference.
Multimodal AI
For most of deep learning's history, models were trained on a single modality: text, images, or audio. A language model understood text but was blind to images. A vision model could classify photographs but could not answer questions about them. Multimodal AI breaks down these barriers, enabling a single model to reason across different types of input.
From Specialist to Generalist
The progression has been rapid. CLIP (Contrastive Language-Image Pre-training, Radford et al., OpenAI 2021) showed that a model trained on 400 million image-text pairs from the internet could learn a shared embedding space for vision and language. This meant you could find images using natural language queries and vice versa, without any task-specific fine-tuning.
GPT-4V (OpenAI 2023) demonstrated that a large language model could accept images as input and reason about them with the same fluency it applied to text. Gemini 1.5 (Google DeepMind 2024) extended this to video, audio, and documents, with a context window of one million tokens capable of processing an entire feature film. The underlying architecture fuses different modalities through shared transformer layers and cross-attention mechanisms.
Vision + Language
Describe image contents, answer questions about photos, read charts and screenshots, generate image captions, and extract structured data from visual documents. Models: GPT-4o, Claude 3.5 Sonnet, Gemini 1.5 Pro.
Audio + Language
Speech-to-text transcription, real-time voice conversation, audio question answering, and speech translation without a separate ASR pipeline. Models: Whisper (OpenAI 2022), Gemini 1.5, GPT-4o audio.
Text to Image / Video
Generate images from text descriptions using diffusion models (DALL-E 3, Stable Diffusion, Midjourney). Video generation from text prompts (Sora, OpenAI 2024; Veo, Google DeepMind 2024).
Diffusion Models
Text-to-image generation is powered by a separate class of model from transformers: diffusion models. The core idea (Sohl-Dickstein et al. 2015; Ho et al. 2020) is to train a neural network to reverse a gradual noising process. Starting from pure noise, the model iteratively denoises the image, guided at each step by a text prompt embedded through a language model. Stable Diffusion (Rombach et al., Stability AI 2022) introduced latent diffusion, which performs the denoising in a compressed latent space rather than pixel space, making generation dramatically faster.
DALL-E 3 (OpenAI 2023) further integrated language understanding by training the image generator to condition on detailed captions generated by GPT-4. This made prompt-following far more reliable and reduced the need for the elaborate prompt engineering that earlier text-to-image systems required.
Multimodal capabilities change what is buildable. A medical imaging assistant that reads both the scan and the clinical notes. A document processing system that understands tables, charts, and prose together. An education platform that answers questions about diagrams and worked examples. If your application involves multiple types of information, multimodal models can often handle the task in a single call rather than requiring separate pipelines stitched together.
AI Agents
A language model in isolation can answer questions and generate text, but it cannot take actions. It cannot search the web for current information, execute code, send an email, interact with a website, or query a database. AI agents extend language models with the ability to use tools and act in the world across a sequence of steps.
What Makes an Agent?
An agent has three components beyond a base language model: tools it can invoke (functions, APIs, browsers, code interpreters), a planning loop that decides which tool to use and when, and memory that preserves state across multiple steps. The ReAct pattern (Yao et al. 2022) interleaves reasoning and acting: at each step, the model reasons about what it knows and what it needs, selects a tool, observes the result, and reasons again. This loop continues until the task is complete.
Function Calling
Modern LLM APIs expose agents through function calling (also called tool use). You define a set of functions with JSON schemas describing their names, parameters, and descriptions. The model decides when to call a function, generates the arguments, and you execute the function in your code and return the result. The model incorporates the result and continues reasoning.
from openai import OpenAI import json client = OpenAI() # Define the tools (functions) available to the model tools = [ { "type": "function", "function": { "name": "get_weather", "description": "Returns current weather for a given city", "parameters": { "type": "object", "properties": { "city": {"type": "string", "description": "City name, e.g. 'London'"} }, "required": ["city"] } } } ] messages = [{"role": "user", "content": "What is the weather like in Accra right now?"}] # First call: model decides to call the tool response = client.chat.completions.create( model="gpt-4o-mini", messages=messages, tools=tools ) tool_call = response.choices[0].message.tool_calls[0] args = json.loads(tool_call.function.arguments) print(f"Model wants to call: {tool_call.function.name}({args})") # Execute the real function and return the result to the model # (in a real system, this would call a weather API) weather_result = '{"temperature": 31, "condition": "Sunny", "humidity": 72}' messages += [response.choices[0].message, {"role": "tool", "tool_call_id": tool_call.id, "content": weather_result}] # Second call: model now answers the user using the tool result final = client.chat.completions.create(model="gpt-4o-mini", messages=messages) print(final.choices[0].message.content)
LangChain (Harrison Chase, 2022) and LlamaIndex are popular Python frameworks for building agent pipelines with pre-built tool integrations, memory backends, and retrieval utilities. The Anthropic Agent SDK, OpenAI Agents SDK, and Microsoft AutoGen (2023) provide higher-level abstractions for multi-agent systems where multiple specialised agents collaborate on a task. Google's Agent Development Kit (ADK) was released in 2025. For production use, these frameworks significantly reduce boilerplate.
Reasoning Models
Standard language models generate responses token by token in a single forward pass. For complex problems requiring multi-step logical deduction, mathematical proof, or careful planning, this is limiting. The model has, in effect, a fixed amount of "thinking time" proportional to the length of the answer.
Reasoning models are trained specifically to think before they answer, generating extensive internal reasoning traces (sometimes called "thinking tokens") before producing a final response. This gives the model time to explore multiple paths, backtrack when an approach fails, and verify its own conclusions.
OpenAI o1 and Beyond
OpenAI's o1 model (released September 2024) was the first widely available reasoning model. It was trained using reinforcement learning to produce internal chain-of-thought reasoning that is longer and more structured than the chain-of-thought prompting technique covered in Lesson 5.3. On the AIME 2024 mathematics competition, o1 scored 74% compared to 13% for GPT-4o. On the International Mathematics Olympiad qualification exam, it scored 83%. These results represent a qualitative jump in formal reasoning capability.
DeepSeek-R1 (DeepSeek AI, January 2025) replicated and in some benchmarks exceeded o1's reasoning performance while being open-weight and at a fraction of the training cost. This was a notable demonstration that state-of-the-art reasoning capability was not restricted to the largest, most resource-rich labs. OpenAI o3 (December 2024 preview) set a new record on ARC-AGI, a benchmark specifically designed to resist pattern matching and require genuine generalisation.
How Reasoning Models Are Trained
The key insight is to reward the quality of the final answer rather than supervising the intermediate steps. The model is trained with reinforcement learning where it receives a reward signal based on whether its final answer is correct. Over many training iterations, the model discovers that producing a long, structured thinking process before answering leads to better final answers and higher rewards. This self-discovered reasoning is in many ways richer than the hand-crafted chain-of-thought prompting strategies described in Lesson 5.3.
Generating long thinking traces increases the number of tokens produced substantially. Reasoning model API calls are significantly more expensive than standard model calls for the same final answer length. For tasks where the problem is well-defined and straightforward, a standard model with a good prompt will often be more cost-effective. Reasoning models are most valuable for tasks that genuinely require multi-step deduction: mathematics, code debugging, logical analysis, and research.
Open vs Closed Models
A significant structural divide has emerged in the AI landscape between closed models accessed only through proprietary APIs and open-weight models whose parameters are publicly released and can be run locally.
Closed / Proprietary
- Accessed via API only; weights not released
- Examples: GPT-4o (OpenAI), Claude 3.5 Sonnet (Anthropic), Gemini Ultra (Google)
- Typically strongest benchmark performance at launch
- No data privacy guarantees; data sent to third-party servers
- No customisation beyond fine-tuning APIs
- Ongoing API costs; risk of model deprecation or price changes
- Alignment and safety work done by provider, not visible to user
Open-Weight
- Weights publicly released; can run locally
- Examples: Llama 3 (Meta), Mistral 7B / Mixtral (Mistral AI), Phi-3 (Microsoft), Gemma (Google), DeepSeek-R1
- Can be fine-tuned on proprietary data without sending data to a third party
- Suitable for air-gapped and regulated environments (healthcare, finance, defence)
- Requires local compute; smaller models trail frontier closed models on benchmarks
- Community ecosystem of fine-tuned variants (Hugging Face Hub)
The Efficiency Frontier
For most of 2023, the gap between the best closed models and the best open models was substantial. That gap has narrowed significantly. Llama 3.1 405B (Meta, July 2024) matched GPT-4 on several benchmarks. Mistral Large 2 competed with GPT-4o on coding tasks. More importantly, smaller open models have improved dramatically: Phi-3-mini (Microsoft 2024, 3.8B parameters) achieved performance comparable to GPT-3.5 on many tasks while fitting on a mobile device.
Parameter-efficient fine-tuning techniques, particularly LoRA (Low-Rank Adaptation, Hu et al. 2021) and its variants, make it possible to adapt open-weight models to specific tasks with modest GPU resources. A 7B model fine-tuned on domain-specific data can outperform a 70B general model on that domain.
For prototyping and tasks without strict data privacy requirements, closed model APIs offer the fastest path to a working system. For production systems in regulated industries, or when fine-tuning on proprietary data is critical, open-weight models running on your own infrastructure are often the better long-term choice. The decision is not purely technical; it involves cost, compliance, and risk tolerance.
AI Safety and Alignment
As AI systems become more capable, the question of whether they will behave as intended across all circumstances becomes increasingly consequential. AI safety is the research field concerned with ensuring that AI systems are reliable, controllable, and aligned with human values even as they become more powerful.
The Alignment Problem
A system is aligned when it reliably pursues the goals its designers intend, including in situations not explicitly anticipated during training. Alignment is hard because it is difficult to fully specify what we want, training objectives are imperfect proxies for our actual goals, and distributional shift means models encounter situations in deployment that were not present in training.
RLHF, covered in Lesson 5.2, is currently the primary alignment technique in production systems. It is effective but has known limitations: the reward model that guides RL training can itself be gamed (reward hacking), and aligning to the preferences of a small pool of labellers may not generalise to all users and contexts.
Current Safety Approaches
| Approach | Description | Limitations |
|---|---|---|
| RLHF | Fine-tune using human preference feedback via a reward model | Reward hacking; labeller subjectivity; expensive at scale |
| Constitutional AI (Anthropic) | Model critiques and revises its own responses using a set of principles | Principles must be carefully specified; model may find loopholes |
| DPO | Direct Preference Optimisation: fine-tune directly from preference pairs without a separate reward model | Still depends on quality of preference data |
| Interpretability | Understand what is happening inside the model (Anthropic's mechanistic interpretability; sparse autoencoders) | Extremely difficult at scale; most circuits not yet understood |
| Red-teaming | Systematic adversarial probing to find failure modes before deployment | Cannot cover all possible inputs; attacker-defender asymmetry |
| Evals | Structured evaluation benchmarks for dangerous capabilities (bioweapons uplift, cyberoffense, deception) | Hard to design comprehensive evals; models improve faster than evals |
Mechanistic Interpretability
A growing research direction called mechanistic interpretability attempts to reverse-engineer the internal workings of neural networks, identifying the circuits (groups of neurons and attention heads) responsible for specific capabilities. Anthropic's research in this area has identified features corresponding to concepts like "the Golden Gate Bridge" or "DNA" in Claude's internal activations, and has found evidence of deliberate planning in the residual stream of transformer models. This work is foundational for building AI systems we can genuinely inspect and verify, rather than only test empirically.
Several frontier AI labs have published responsible scaling policies (Anthropic) or preparedness frameworks (OpenAI) that define thresholds of dangerous capability. The policies commit to pausing development or deployment if models cross certain thresholds on assessments of autonomy, cyberoffense capability, biological and chemical weapons uplift, and deception. These frameworks represent an attempt to build safety gates into the development process rather than evaluating safety only at deployment time.
Economic and Societal Impact
The economic effects of AI are already measurable and will intensify. Understanding them is important both for making career decisions and for thinking about the policy frameworks needed to manage the transition well.
Productivity and Labour
Studies of AI tools in professional settings show consistent productivity gains for workers who adopt them. A GitHub Copilot study (Peng et al. 2023) found that developers using AI code completion completed tasks 55% faster. A study of customer service agents using AI assistance (Brynjolfsson, Li, and Raymond 2023) found a 14% increase in resolutions per hour, with the largest gains for less experienced workers. A randomised trial of consultants using GPT-4 (Noy and Zhang, 2023) found AI users completed tasks 25% faster with 40% higher quality ratings.
These productivity gains are real, but they are not evenly distributed. The tasks most easily automated tend to be the structured, routine parts of knowledge work: first drafts of reports, boilerplate code, straightforward customer queries, and data extraction. The tasks that remain require judgment, relationship-building, ethical reasoning, novel problem-solving, and contextual expertise. This shifts the premium from information access and execution speed toward these harder-to-automate qualities.
The Education and Credentialling Challenge
AI significantly changes what skill gaps matter and what skills are worth developing. The ability to use AI tools effectively is becoming as fundamental as the ability to use a search engine. At the same time, the deep domain knowledge that allows you to evaluate AI outputs, catch errors, and direct AI effectively remains valuable precisely because it is not easily automated. The people best positioned in an AI-integrated world are those who combine domain expertise with practical AI literacy.
Geopolitical Dimensions
AI capability has become a matter of national strategic interest. The United States, China, and the European Union are all pursuing distinct approaches to AI development and governance. US export controls on advanced semiconductors (Nvidia A100 and H100 chips, and subsequent generations) are an attempt to maintain an advantage in AI training capacity. China's domestic AI investment includes state-sponsored labs and significant open-weight contributions (DeepSeek, Qwen). The EU AI Act represents the first comprehensive regulatory framework for AI systems globally, likely to influence AI governance internationally.
Amidst the rapid change, some things remain constant. Good problem formulation, intellectual honesty, and the ability to communicate clearly are more valuable than ever. Understanding why a model produces a certain output, not just how to get it to produce one, is a durable advantage. And building things that are actually useful to real people, rather than technically impressive without practical value, remains the measure of whether the work matters.
What You Should Do Next
You have now completed the entire AI Crash Course. You have built a conceptual map of the field from the mathematics of a single neuron through to the frontier of reasoning models and AI agents. The question now is how to keep building on this foundation.
Key Takeaways
- Multimodal AI integrates text, images, audio, and video in unified models. CLIP, GPT-4V, and Gemini represent a shift from specialist to generalist systems. Diffusion models (Stable Diffusion, DALL-E 3) power text-to-image generation through iterative denoising.
- AI agents extend language models with tools, planning loops, and memory, enabling multi-step task completion. Function calling (tool use) is the API mechanism that makes agents practical. Frameworks like LangChain, LlamaIndex, and AutoGen reduce the boilerplate of building agent systems.
- Reasoning models (OpenAI o1, o3; DeepSeek-R1) are trained with reinforcement learning to generate extensive internal thinking before answering. They show qualitative improvements on formal reasoning tasks but cost more per call.
- Open-weight models (Llama 3, Mistral, Phi-3) have closed the gap with frontier closed models significantly. LoRA fine-tuning makes domain adaptation practical on consumer hardware. The right choice between open and closed models depends on privacy requirements, regulatory constraints, and resource availability.
- AI alignment remains an unsolved research problem. Current approaches (RLHF, Constitutional AI, DPO, red-teaming, interpretability) each address part of the problem. Mechanistic interpretability is a promising long-term direction for understanding model behaviour from first principles.
- Measured productivity gains from AI in professional settings are real and consistent across studies. The premium shifts toward judgment, domain expertise, and the ability to evaluate and direct AI outputs effectively.
- The best position in an AI-integrated world combines deep domain knowledge with practical AI literacy. Build things, engage with research, and specialise deliberately.
Going Deeper
Want to think seriously about where AI is heading? Start here.
Reflect
Before you move on
No right answers here. These questions are for you.
Throughout history, new technologies have disrupted old jobs and created new ones. Is AI different from the printing press or the industrial revolution, and if so, how?
Previous disruptions replaced physical labour or narrow clerical tasks. AI can now perform cognitive work that spans language, reasoning, and creativity. The speed of displacement may also be faster than past transitions, leaving less time for workers and institutions to adapt. That said, humans have consistently found new roles after prior transitions. Whether AI follows the same pattern is one of the most important open questions of this era.
Multiple AI labs say they believe they are building potentially dangerous technology but press on anyway, citing the logic that "if we don't, someone else will." Is that reasoning sound?
This is one of the central tensions in AI development. The argument has some strategic validity if unregulated actors would indeed proceed without safety investment. But it can also become a self-fulfilling justification for any level of risk. It also assumes that being first confers enough control to mitigate harm, which is unproven. Recognising the logic as a pressure on behaviour, rather than an airtight justification, is an important step in thinking critically about how the field operates.
You have just completed this course. What is one concrete thing you want to build, change, or question because of what you have learned?
This is yours to answer. There is no model response here. The point of the course was not to make you feel informed about AI in the abstract; it was to give you enough understanding to act on it in your own context. Write the answer down somewhere. Come back to it in six months.